Papers with German dataset

5 papers
ADEA: An Argumentative Dialogue Dataset on Ethical Issues Concerning Future A.I. Applications (2024.lrec-main)

Copied to clipboard

Challenge: Introducing ADEA: a dataset that captures online dialogues and focuses on ethical issues related to future AI applications.
Approach: They propose a German dataset that captures online dialogues on ethical issues . the dataset includes over 2800 labeled user utterances on four different topics . they use an argument graph as the system's knowledge base and an annotation scheme .
Outcome: The proposed dataset includes over 2800 user utterances on four ethical topics . the aim is to improve knowledge about AI ethics topics through argumentative dialogues .
Our kind of people? Detecting populist references in political debates (2023.findings-eacl)

Copied to clipboard

Challenge: Existing literature on populism has only limited agreement on its exact properties .
Approach: They propose a cross-lingual dataset to identify populist rhetoric in text . they propose 'hierarchical' annotation procedure to annotate populist references .
Outcome: The proposed dataset can be used to investigate how political actors talk about The Elite and The People and to study how populist rhetoric is used as a strategic device.
GLoHBCD: A Naturalistic German Dataset for Language of Health Behaviour Change on Online Support Forums (2022.lrec-1)

Copied to clipboard

Challenge: Existing motivational interviewing methods lack the deep understanding of user utterances that is essential to the spirit of motivational interviews.
Approach: They propose to use a German dataset of naturalistic language around health behaviour change to examine the motivational state of the user.
Outcome: The proposed dataset of naturalistic language around health behaviour change is based on a weight loss forum in germany and is evaluated using theoretically grounded motivational interviewing categories.
On the Impact of Cross-Domain Data on German Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Traditionally, large language models have been trained on general web crawls or domain-specific data.
Approach: They present a German dataset and a dataset aimed at containing high-quality data to examine the importance of data diversity over quality.
Outcome: The proposed model outperforms models trained on quality data on multiple downstream tasks.
Using Pre-Trained Language Models in an End-to-End Pipeline for Antithesis Detection (2024.lrec-main)

Copied to clipboard

Challenge: Rhetorical figures are a "departure from the normal usage" of language . features of metaphors, irony and sarcasm enhance performance of several NLP tasks.
Approach: They propose a pipeline approach to detect rhetorical figures using large language models by splitting text into phrases and identifying parallel phrases with a syntactically parallel structure.
Outcome: The proposed approach outperforms state-of-the-art methods by an F1 score of 65.11 %.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations